Research

The project shifted from evaluating a particular model towards evaluating a reproducible, model-neutral method.

The pilot clarified an important conceptual point: the AI platform is not the method. It is the implementation environment in which the method is executed. A study framed only as ‘ChatGPT versus reviewers’ would age quickly and would be difficult to reproduce when the interface, model or product changed. The more durable intervention is the structured protocol: the source restrictions, locked assessment target, sequential instructions, response rules, verification loop, version control and human oversight process.A study framed only as ‘ChatGPT versus reviewers’ would age quickly and would be difficult to reproduce when the interface, model or product changed. The more durable intervention is the structured protocol: the source restrictions, locked assessment target, sequential instructions, response rules, verification loop, version control and human oversight process.

This does not make the model irrelevant. The exact model, interface, date and prompt version must still be logged because performance may change. It does, however, change the scientific objective. The project is not trying to build another proprietary risk-of-bias tool. It is testing whether available generative AI systems can execute a transparent appraisal protocol reliably enough to support reviewers.

That framing also permits future replication across models and tasks. A later system can be tested using the same locked source package and parameterisation, while any performance change can be attributed more clearly to the implementation rather than an undocumented alteration of the method.In brief: the model may be replaced; the rules must be inspectable.